Predicting reliable regions in protein sequence alignments

نویسندگان

  • Melissa S. Cline
  • Richard Hughey
  • Kevin Karplus
چکیده

MOTIVATION Protein sequence alignments have a myriad of applications in bioinformatics, including secondary and tertiary structure prediction, homology modeling, and phylogeny. Unfortunately, all alignment methods make mistakes, and mistakes in alignments often yield mistakes in their application. Thus, a method to identify and remove suspect alignment positions could benefit many areas in protein sequence analysis. RESULTS We tested four predictors of alignment position reliability, including near-optimal alignment information, column score, and secondary structural information. We validated each predictor against a large library of alignments, removing positions predicted as unreliable. Near-optimal alignment information was the best predictor, removing 70% of the substantially-misaligned positions and 58% of the over-aligned positions, while retaining 86% of those aligned accurately.

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Combining sequence and structure information in protein alignments

For distantly related proteins, alignmentsbased on structural information are more reliable than traditional sequence alignments. However, when structural comparison leaves some ambiguity in alignment, sequence information can provide valuable additional information to discriminate between multiple alternatives. In this paper we present a Bayesianmodel that incorporates sequence information int...

متن کامل

Tracking repeats using significance and transitivity

MOTIVATION Internal repeats in coding sequences correspond to structural and functional units of proteins. Moreover, duplication of fragments of coding sequences is known to be a mechanism to facilitate evolution. Identification of repeats is crucial to shed light on the function and structure of proteins, and explain their evolutionary past. The task is difficult because during the course of e...

متن کامل

GenTHREADER: an efficient and reliable protein fold recognition method for genomic sequences.

A new protein fold recognition method is described which is both fast and reliable. The method uses a traditional sequence alignment algorithm to generate alignments which are then evaluated by a method derived from threading techniques. As a final step, each threaded model is evaluated by a neural network in order to produce a single measure of confidence in the proposed prediction. The speed ...

متن کامل

RASCAL: Rapid Scanning and Correction of Multiple Sequence Alignments

MOTIVATION Most multiple sequence alignment programs use heuristics that sometimes introduce errors into the alignment. The most commonly used methods to correct these errors use iterative techniques to maximize an objective function. We present here an alternative, knowledge-based approach that combines a number of recently developed methods into a two-step refinement process. The alignment is...

متن کامل

Ca Bios Invited Review

The problem of predicting protein structure from the sequence remains fundamentally unsolved despite more than three decades of intensive research effort. However, new and promising methods in three-dimensional (3D), 2D and ID prediction have reopened the field. Mean-forcepotentials derived from the protein databases can distinguish between correct and incorrect models (3D). Inter-residue conta...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

عنوان ژورنال:
  • Bioinformatics

دوره 18 2  شماره 

صفحات  -

تاریخ انتشار 2002